Papers with alignment challenge

2 papers
Comparing human and LLM politeness strategies in free production (2025.emnlp-main)

Copied to clipboard

Challenge: Polite speech poses a fundamental alignment challenge for large language models (LLMs).
Approach: They compare human and LLM responses to English-language scenarios to determine whether they employ a similarly context-sensitive repertoire.
Outcome: The results show that large models replicate key effects from the computational pragmatics literature and human evaluators prefer LLM-generated responses in open-ended contexts.
BiasGRPO: Stabilizing Bias Mitigation in High-Variance Reward Landscapes via Group-Relative Policy Optimization (2026.findings-acl)

Copied to clipboard

Challenge: Recent preference-based fine-tuning methods have limited exploration in offline training . previous methods have been limited by the lack of exploration inherent in offline learning .
Approach: They propose a method that normalizes rewards across a group of completed tasks to mitigate social bias in Large Language Models.
Outcome: The proposed approach outperforms DPO and PPO in multiple benchmarks . it can overcome limitations of previous preference-based methods .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations